Papers with gradient computation
Parameter-efficient Tuning for Large Language Model without Calculating Its Gradients (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent parameter-efficient tuning methods can only save 30% of training memory . gradient computation and backpropagation are still necessary for these methods . |
| Approach: | They propose a parameter-efficient tuning method that can be used to fine-tune large language models without calculating gradients. |
| Outcome: | The proposed method saves 30% of training memory and improves performance on large language models. |
Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to identify key neurons for interpretability of multi-modal large language models are unclear. |
| Approach: | They propose a method to identify key neurons for interpretability by multi-modal large language models. |
| Outcome: | The proposed method improves conventional works upon efficiency and applied range by removing needs of costly gradient computation. |
Reliable Gradient-free and Likelihood-free Prompt Tuning (2023.findings-eacl)
Copied to clipboard
| Challenge: | Large pre-trained language models are often offered as black-box APIs due to privacy or commercial constraints. |
| Approach: | They propose to tune the soft prompts without requiring gradient computation and extend the model to include a distribution over prompts. |
| Outcome: | The proposed methods are competitive with gradient-based approaches with full access to the PLM. |
AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for making decisions in grounded environments require costly gradient computation or lengthy in-context demonstrations. |
| Approach: | They propose an approach to guide LLM-based agents to accomplish interactive decision-making tasks by using an LLM prompt and a task-solving plan. |
| Outcome: | The proposed approach outperforms human-written demonstrations on ALFWorld and HotpotQA by 8%. |
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search (2026.findings-acl)
Copied to clipboard
Yifei Zhang, Xu Yang, Xiao Yang, Bowen Xian, Qizheng Li, Shikai Fang, Jingyuan Li, Jian Wang, Minrui Xu, Yuge Zhang, Weiqing Liu, Jiang Bian
| Challenge: | LLM-based agents for machine learning engineering rely on tree search to rank candidates. |
| Approach: | They propose an LLM-based agent that operationalizes gradient-based optimization. |
| Outcome: | The proposed agent achieves a state-of-the-art 35.1% any-medal rate on MLE-Bench with a limited budget on a single GPU. |
Full Parameter Fine-tuning for Large Language Models with Limited Resources (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) require massive GPU resources for training. |
| Approach: | They propose a parameter-efficient optimization that fuses the gradient computation and parameter update in one step to reduce memory usage. |
| Outcome: | The proposed method reduces memory usage to 10.8% compared to the standard approach. |
Discretized Integrated Gradients for Explaining Language Models (2021.emnlp-main)
Copied to clipboard
| Challenge: | Integrated Gradients (IG) is widely adopted due to its desirable explanation axioms and the ease of gradient computation. |
| Approach: | They propose an attribution-based explanation algorithm that uses averaging the model's output gradient interpolated along a straight-line path in the input data space. |
| Outcome: | The proposed method is compared with IG on multiple sentiment classification datasets. |
Memory-Efficient Fine-Tuning of Transformers via Token Selection (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning require caching of intermediate activations to update weights during the backward pass. |
| Approach: | They develop a method to reduce memory usage in fine-tuning of transformers by backpropagating through just a subset of input tokens. |
| Outcome: | The proposed method reduces memory usage and memory footprint on large transformer models . it can be easily combined with existing methods like LoRA, reducing memory cost . |